Papers with Bangla language
BanglaBook: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing literature on Bangla Sentiment Analysis (SA) has limited data and cross-domain adaptability. |
| Approach: | They present a large-scale dataset of Bangla book reviews with 158,065 samples . they employ a range of machine learning models to establish baselines including SVM, LSTM, and Bangla-BERT. |
| Outcome: | The proposed model improves performance over models that rely on manual features. |
BanStereoSet: A Dataset to Measure Stereotypical Social Biases in LLMs for Bangla (2025.findings-acl)
Copied to clipboard
| Challenge: | ***BanStereoSet*** is a dataset designed to evaluate stereotypical social biases in multilingual LLMs for the Bangla language. |
| Approach: | They propose to localize the content from StereoSet, IndiBias, and kamruzzaman-etal's datasets to capture biases prevalent within the Bangla language. |
| Outcome: | The proposed dataset consists of 1,194 sentences spanning 9 categories of bias: race, profession, gender, ageism, beauty, beauty in profession, region, caste, and religion. |
Customizing Grapheme-to-Phoneme System for Non-Trivial Transcription Problems in Bangla Language (N19-1)
Copied to clipboard
Sudipta Saha Shubha, Nafis Sadeq, Shafayat Ahmed, Md. Nahidul Islam, Muhammad Abdullah Adnan, Md. Yasin Ali Khan, Mohammad Zuberul Islam
| Challenge: | Existing methods for Grapheme to phoneme conversion in Bangla language are mostly rule-based. |
| Approach: | They propose to use a lexicon to train a robust Grapheme to phoneme conversion system in Bangla language. |
| Outcome: | The proposed method outperforms other state-of-the-art approaches for G2P conversion in Bangla language. |
BanNERD: A Benchmark Dataset and Context-Driven Approach for Bangla Named Entity Recognition (2025.findings-naacl)
Copied to clipboard
Md. Motahar Mahtab, Faisal Ahamed Khan, Md. Ekramul Islam, Md. Shahad Mahmud Chowdhury, Labib Imam Chowdhury, Sadia Afrin, Hazrat Ali, Mohammad Mamun Or Rashid, Nabeel Mohammed, Mohammad Ruhul Amin
| Challenge: | In a cross-dataset evaluation, models trained on BanNERD consistently outperformed those trained on four existing Bangla NER datasets. |
| Approach: | They propose to use Bangla as a language to create the most extensive human-annotated and validated Bangla NLP dataset. |
| Outcome: | The proposed method outperforms existing methods on Bangla NER datasets and performs competitively on English datasets. |
Preparation of Bangla Speech Corpus from Publicly Available Audio & Text (2020.lrec-1)
Copied to clipboard
Shafayat Ahmed, Nafis Sadeq, Sudipta Saha Shubha, Md. Nahidul Islam, Muhammad Abdullah Adnan, Mohammad Zuberul Islam
| Challenge: | Automated speech recognition systems require large annotated speech corpus for training. |
| Approach: | They propose to use publicly available Bangla audiobooks and TV news recordings as input to prepare a large speech corpus with reasonable confidence. |
| Outcome: | The proposed algorithm outperforms the existing speech corpus and the existing corpus with speaker diarization and gender detection. |